Scalability of a Distributed Neural Information Retrieval System
نویسندگان
چکیده
AURA (Advanced Uncertain Reasoning Architecture) is a generic family of techniques and implementations intended for high-speed approximate search and match operations on large unstructured datasets [1]. AURA technology is fast, economical, and offers unique advantages for finding near-matches not available with other methods. AURA is based upon a high-performance binary neural network called a Correlation Matrix Memory (CMM). Typically, several CMM elements are used in combination to solve soft or fuzzy pattern-matching problems. AURA takes large volumes of data and constructs a special type of compressed index. AURA finds exact and near-matches between indexed records and a given query, where the query itself may have omissions and errors. The degree of nearness required during matching can be varied through thresholding techniques. The PCI-based PRESENCE (PaRallEl Structured Neural Computing Engine) card is a hardware-accelerator architecture for the core CMM computations needed in AURA-based applications. The card is designed for use in low-cost workstations and incorporates 128MByte of lowcost DRAM for CMM storage. The E-Science project, Distributed Aircraft Maintenance Environment (DAME), will use AURA technology to process hundreds of gigabytes of aircraft aeroengine diagnostic information. The size of this data means that we are unable to map it to a locally implemented software or hardware CMMs. Therefore, we need to make use of distributed AURA methods that allow a CMM to be striped over multiple PRESENCE cards over a cluster. It is important therefore, that we determine how the performance of the distributed AURA system scales with increasing dataset size. To investigate the scalability of the distributed AURA system, we implement a word-to-document index of an AURA-based information retrieval system, called MinerTaur[2], over a distributed-PRESENCE CMM. This follows on from a previous paper that compared local and distributed AURA performance [3]. Here we also give updated performance figures for the distributed AURA system, which has since been substantially improved. MinerTaur comprises three modules, a spell checking pre-processor to identify errors in the user query, a synonym hierarchy to allow paraphrased documents to be matched and a word-document indexing module to identify documents matching particular query words. All modules rely on a fast efficient data structure to under-pin the system. To provide this we implement MinerTaur using binary Correlation Memory Matrices (CMMs) on PCI-based PRESENCE cards in a Beowulf PC cluster named Cortex-1. Cortex-1 consists of seven 500MHz PC nodes connected by 100Mbit Ethernet, six nodes of which contain 28 PCI-PRESENCE cards. The document corpus used is 571 MBytes in size and contains 476,672 Reuters document abstracts [4]. 62,903 keywords were extracted from the file to index the documents. Mapping this data onto the CMM with a single bit vector representing both word and document, this will fill the whole weights memory available with the Cortex-1 cluster’s 28-cards.
منابع مشابه
Dynamic configuration and collaborative scheduling in supply chains based on scalable multi-agent architecture
Due to diversified and frequently changing demands from customers, technological advances and global competition, manufacturers rely on collaboration with their business partners to share costs, risks and expertise. How to take advantage of advancement of technologies to effectively support operations and create competitive advantage is critical for manufacturers to survive. To respond to these...
متن کاملA Radon-based Convolutional Neural Network for Medical Image Retrieval
Image classification and retrieval systems have gained more attention because of easier access to high-tech medical imaging. However, the lack of availability of large-scaled balanced labelled data in medicine is still a challenge. Simplicity, practicality, efficiency, and effectiveness are the main targets in medical domain. To achieve these goals, Radon transformation, which is a well-known t...
متن کاملHadoop Scalability and Performance Testing in Heterogeneous Clusters
This paper aims to evaluate cluster configurations using Hadoop in order to check parallelization performance and scalability in information retrieval. This evaluation will establish the necessary capabilities that should be taken into account specifically on a Distributed File System (HDFS: Hadoop Distributed File System), from the perspective of storage and indexing techniques, and queriy dis...
متن کاملPerformance Comparison of Clustered and Replicated Information Retrieval Systems
The amount of information available over the Internet is increasing daily as well as the importance and magnitude of Web search engines. Systems based on a single centralised index present several problems (such as lack of scalability), which lead to the use of distributed information retrieval systems to effectively search for and locate the required information. A distributed retrieval system...
متن کاملTowards a Scalable Networked Retrieval System for Searching Multimedia Databases
In this paper the architecture of a distributed and scalable multimedia information retrieval system (Dsmily) is described. The system consists of hierarchically organized networked nodes and is designed to integrate existing dynamic multimedia databases. The document ranking process as well as the preselection of databases to be searched, both tasks are based on a probabilistic model for distr...
متن کاملAttribute-based Access Control for Cloud-based Electronic Health Record (EHR) Systems
Electronic health record (EHR) system facilitates integrating patients' medical information and improves service productivity. However, user access to patient data in a privacy-preserving manner is still challenging problem. Many studies concerned with security and privacy in EHR systems. Rezaeibagha and Mu [1] have proposed a hybrid architecture for privacy-preserving accessing patient records...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2002